Papers with ASR model

10 papers
STT4SG-350: A Speech Corpus for All Swiss German Dialect Regions (2023.acl-short)

Copied to clipboard

Challenge: We present a corpus of Swiss German speech annotated with Standard German text at the sentence level.
Approach: They present a corpus of Swiss German speech annotated with Standard German sentences . they use a web app to show the speakers standard German sentences and record them .
Outcome: The corpus contains 343 hours of speech from all Swiss German dialect regions . it is the largest public speech corpus for Swiss German to date .
RED-ACE: Robust Error Detection for ASR using Confidence Embeddings (2022.emnlp-main)

Copied to clipboard

Challenge: ASR Error Detection (AED) models post-process the output of Automatic Speech Recognition systems, in order to detect transcription errors.
Approach: They propose to use ASR model's word-level confidence scores to combine ASR models with transcribed text to improve AED performance.
Outcome: The proposed models combine the confidence scores and transcribed text into a contextualized representation.
End-to-End Speech Recognition and Disfluency Removal (2020.findings-emnlp)

Copied to clipboard

Challenge: Disfluency detection is usually an intermediate step between an automatic speech recognition system and a downstream task.
Approach: They propose to train models to directly map disfluent speech into fluent transcripts without relying on a separate disfluency detection model.
Outcome: The proposed models learn to generate fluent transcripts, but their performance is slightly worse than a baseline pipeline approach consisting of an ASR system and a specialized disfluency detection model.
DITTO: Data-efficient and Fair Targeted Subset Selection for ASR Accent Adaptation (2023.acl-long)

Copied to clipboard

Challenge: State-of-the-art automatic speech recognition systems exhibit disparate performance on varying speech accents.
Approach: They propose to use submodular mutual information to find the most informative set of utterances matching a target accent within a fixed budget.
Outcome: The proposed model is 3-5 times more label-efficient on the Indic-TTS and L2 datasets than other methods.
Error-preserving Automatic Speech Recognition of Young English Learners’ Language (2024.acl-long)

Copied to clipboard

Challenge: State-of-the-art speech recognition models are often trained on adult read-aloud data by native speakers and do not transfer well to young language learners’ speech.
Approach: They propose to use an automated speech recognition module to train language learners' speaking skills on spontaneous speech by young language learners.
Outcome: The proposed model improves on 85 hours of English audio spoken by Swiss learners and preserves their mistakes.
Beyond Common Words: Enhancing ASR Cross-Lingual Proper Noun Recognition Using Large Language Models (2024.findings-emnlp)

Copied to clipboard

Challenge: In this work, we address the challenge of cross-lingual proper noun recognition in automatic speech recognition systems where proper nodes in an utterance may originate from a language different from the language in which the ASR system is trained.
Approach: They propose a dictionary-based method to correct ASR predictions in a large language model .
Outcome: The proposed method significantly reduces word error rates across cross-lingual proper noun recognition tasks involving three secondary languages.
MuPe Life Stories Dataset: Spontaneous Speech in Brazilian Portuguese with a Case Study Evaluation on ASR Bias against Speakers Groups and Topic Modeling (2025.coling-main)

Copied to clipboard

Challenge: Recent datasets for automatic speech recognition in Brazilian Portuguese lack diversity in terms of age groups, regional accents, and education levels.
Approach: They propose to use a dataset to analyze the impact of ASR in Brazilian Portuguese (BP) they demonstrate that current models are biased regarding age, education, and regional accents.
Outcome: The proposed dataset helps mitigate biases in current ASR models regarding education levels and age groups.
Sequential Randomized Smoothing for Adversarially Robust Speech Recognition (2021.emnlp-main)

Copied to clipboard

Challenge: Existing, naive defenses against adversarial attacks are lagging . a new paper aims to break these defenses with adaptive noise ensembling .
Approach: They propose a randomized smoothing paradigm that can be used to break adversarial attacks . they use speech enhancement methods and a novel use for ASR output ensembling methods .
Outcome: The proposed model is robust to all attacks that use inaudible noise and can only be broken with very high distortion.
Automatic Speech Recognition Datasets in Cantonese: A Survey and New Dataset (2022.lrec-1)

Copied to clipboard

Challenge: In this paper, we address the problem of data scarcity for the Hong Kong Cantonese language . due to the popularization of deep learning, ASR technology has led to a significant improvement in recognizing many languages.
Approach: They propose to use a dataset to analyze the data available for the Hong Kong Cantonese language . they use zh-HK as a source and a state-of-the-art ASR model to build a powerful model .
Outcome: The proposed model improves on the biggest existing dataset, Common Voice zh-HK.
Killkan: The Automatic Speech Recognition Dataset for Kichwa with Morphosyntactic Information (2024.lrec-main)

Copied to clipboard

Challenge: Existing datasets for automatic speech recognition (ASR) in the endangered Kichwa language have been limited.
Approach: They present Killkan, the first dataset for automatic speech recognition (ASR) in the Kichwa language, an indigenous language of Ecuador.
Outcome: The proposed dataset shows that it can be used to build an automatic speech recognition system for the endangered language with reliable quality despite its small size.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations